Goto

Collaborating Authors

 evaluation crisis


Can we fix AI's evaluation crisis?

MIT Technology Review

So far, the way we've tried to answer that question is through benchmarks. These give models a fixed set of questions to answer and grade them on how many they get right. But just like exams like the SAT (an admissions test used by many US colleges), these benchmarks don't always reflect deeper abilities. Lately it feels as if a new AI model drops every week, and every time a company launches one, it comes with fresh scores showing it beating the capabilities of predecessors. On paper, everything appears to be getting better all the time.